1 results listed
Social media platforms such as Twitter have grown at
a tremendous pace in recent years and have become an important
source of data providing information countless field. This situation
was of interest to researchers and many studies on machine
learning and natural language processing were conducted on
social media data. However, the language used in social media
contains a very high amount of noisy data than the formal writing
language. In this article, we present a study on diacritic restoration
which is one of the important difficulties of social media text
normalization in order to reduce the noise problem. Diacritic is a
set of marks used to change the sound values of letters and is used
on many languages besides Turkish. We suggest a 3-step model for
this study to overcome the top of the diacritic restoration problem.
In the first stage, a candidate word producer produces possible
word forms, in the second stage the language validator chooses the
correct word forms and at the final word2vec is used to create
vector representations of the words and make the most
appropriate word choice by using cosine similarities. The
proposed method was tested on both synthetic and real data sets,
and we achieved a relative error reduction of 37.8% in our data
sets compared to the previous study with an average of 94.5%
performance.
International Conference on Advanced Technologies, Computer Engineering and Science
ICATCES
Zeynep Ozer
İlyas özer
Oğuz Findik